Cell Genomics
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Cell Genomics's content profile, based on 172 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.
Ahmed, O. Y.; Saravanan, N.; Rovsing, A. B.; Simpson, D.; Devarajan, A.; Gunn, S.; Singh, T.; Lappalainen, T.; Sanjana, N. E.
Show abstract
Over the past two decades, genome-wide association studies (GWAS) have identified thousands of trait- and disease-associated loci. However, the mechanistic understanding of these loci remains incomplete, which limits our ability to understand gene regulation and cellular programs underlying complex traits, predict disease risk, and develop therapeutics targeted to root causes. Here, we describe the current challenges for using GWAS to prioritize variants for functional follow-up experiments. These challenges span multiple domains, including limitations in data sharing and harmonization, limitations of statistical and functional fine-mapping, and the ambiguity in the added value of emerging deep learning frameworks for variant effect prediction as a complementary approach alongside traditional statistical genetics methods. We analyze these variant prioritization methods and suggest a multi-modal approach for resolving GWAS loci to a focused set of high-confidence variants for functional exploration. Fully realizing the potential of GWAS will require harmonized summary statistics and broader sharing of in-sample linkage disequilibrium (LD) data to enable robust and scalable causal variant prioritization.
Jia, Q.; Lam, M.; Wang, L.; Zhao, F.; Tang, H.; Sarashetti, P.; Li, Z.; Wong, E.; SG10K_Health Consortium, ; Tan, P.; Sim, X.; Ngeow, J.; Lee, J.; Cheng, C.-Y.; Chee, M. L.; Lim, W. K.; Chin, C. W. L.; Karnani, N.; Chong, Y. S.; Sim, W. C.; Lim, C. W.; Bertin, N.; Liu, J.
Show abstract
Tandem repeats (TRs) are implicated in over 70 Mendelian disorders and likely contribute to the "missing heritability" of complex traits and diseases, yet TR variations in Asian populations remain poorly characterized. Here, we constructed an Asian-specific SG10K-TR catalog by leveraging the SG10K_Health Dataset, comprising 916,274 autosomal TR loci genotyped in 9,490 individuals of Chinese (5,528), Malay (1,824), Indian (2,108), and other ancestries (30). Using a novel integrative measure for both repeat length and frequency variations, TRDDS, we found that population-level TR variations are selectively constrained in coding and promoter regions, whereas the enrichment of TRs with high population diversity was observed in regulatory sites with low chromatin accessibility and pathways related to neuronal functions. We also identified candidate TRs under selection that predominantly targets neuronal and synaptic architecture. Analysis of linkage disequilibrium (LD) patterns revealed that TRs are often poorly tagged by small variants, although we identified 123 candidate functional TRs that may underlie association signals previously attributed to nearby noncoding SNPs. Finally, TR-based GWAS of six anthropometric and lipid traits identified ten loci with genome-wide significant associations, including two novel loci for BMI (LINC02817) and height (UNC45B), and a TR variant as causal candidate for a known GWAS locus at HMGCR for LDL. Together, this study establishes a critical Asian-specific TR resource and highlights the fundamental role of TR diversity in driving evolutionary neuroplasticity and shaping the genetic architecture of complex traits.
Cadiou, S.; Konig, E.; Mapelli, A.; Pontali, G.; Ghasemi-Semeskandeh, D.; Filosi, M.; Ferolito, B. R.; Massi, M. C.; Cuccuru, G.; Jiang, X.; Winicki, G.; Navarro-Gallinad, A.; Gravel-Pucillo, K.; Pirastu, N.; Landini, A.; Sharapov, S.; Rainer, J.; Gogele, M.; Lundin, R.; Mascalzoni, D.; Biasiotto, R.; Blankenburg, H.; De Grandi, A.; Egger, C.; Arend, L.; Woller, F.; Ieva, F.; Soranzo, N.; Cho, K.; Gaziano, J. M.; Zuccolo, L.; Domingues, F. S.; Pattaro, C.; Danesh, J.; Pramstaller, P. P.; Pereira, A.; Di Angelantonio, E.; Fuchsberger, C.; Giambartolomei, C.; Butterworth, A. S.
Show abstract
Circulating plasma proteins are key biomarkers and therapeutic targets, now measurable at scale through high-throughput technologies, yet whether expanding proteomics platforms beyond the classical plasma secretome enhances genetic discovery and causal inference remains poorly understood. Here, we use an expanded SomaScan 7k platform to map the genetic architecture of a broader segment of the plasma proteome and to evaluate how proteome expansion affects pQTL discovery, causal inference and therapeutic target prioritisation. After quality control, we analysed 7,144 aptamers targeting 6,267 proteins in the harmonised dataset of two European cohorts: INTERVAL (n = 9,251 participants) and CHRIS (n = 4,194), and conducted genome-wide pQTL association analyses followed by meta-analysis. We identified 7,870 significant pQTLs (P-value < 1.26 x 10E-11; 1,784 cis, 6,086 trans), of which 2,704 (34%) associations were not reported in five prior large-scale pQTL studies. Newly assessed proteins, which accounted for 53% (1,422/2,704) of the novel associations, were less likely to harbour cis-pQTLs associations (15%) than those in the previous platform version (28%), consistent with their lower expected plasma concentrations and predominantly intracellular localisation. Colocalization analyses revealed widespread sharing of genetic signals across proteins and characterised 22 pleiotropic trans-regulatory hotspots accounting for 68% of all trans-pQTLs. Through two-sample Mendelian randomization analyses on 2,003 phenotypes from the Million Veteran Program, UK Biobank, and FinnGen (combined N > 1.2 million), we identified 6,340 genetically supported protein-trait associations, highlighting disease mechanisms and potential therapeutic opportunities beyond currently drug-targeted circulating proteins. Together, these findings provide a systematic view of the genetic architecture of the expanded plasma proteome and demonstrate that plasma proteome expansion reveals genetically anchored disease biology beyond the classical secretome, while exposing inherent biological and technical constraints of studying low-abundance intracellular proteins in circulation.
Upmeier zu Belzen, J.; Arnoldt, L.; Hollmann, N.; Herrmann, L.; Nguyen, K. M.; Eckhoff, L.; Kohleick, L.; Abou Ghaloun, S.; Schmidt, H.; Hegselmann, S.; Theis, F. J.; Buergel, T.; Steinfeldt, J.; Wild, B.; Eils, R.
Show abstract
Genetic prediction of complex phenotypes typically relies on additive linear models, which scale well but cannot capture non-additive effects or deeply integrate molecular and clinical data. Domain-specific neural networks have driven advances in images, text, and other modalities, but genome-scale neural networks remain challenging because genotypes are sparse and high-dimensional, effective sample sizes are limited, and generic architectures lack interpretability. Here, we introduce the omnigenic neural network, a biologically structured architecture inspired by the omnigenic model of complex traits. The model learns hierarchical representations of biological processes, accommodates multimodal inputs, supports transfer learning, and enables multitask prediction. Models trained in the UK Biobank and evaluated in the All of Us cohort for ischemic heart disease, type 2 diabetes, and schizophrenia outperformed published PGS Catalog and PRS-CSx scores. A multitask model trained across 36 cardiovascular endpoints further outperformed corresponding single-phenotype models and baselines. The architecture provides systems-level interpretability by quantifying the contributions of biological processes, which were consistent with established disease mechanisms. It also captures non-linear interactions between variants. Analysis of these interactions using Integrated Hessians revealed patterns concordant with previously reported epistatic associations. Together, these findings establish the omnigenic neural network as a flexible framework for interpretable, multimodal, and multitask genomic prediction.
Chen, T.; Li, X.; Mazumder, R.; Zhang, H.; Lin, X.
Show abstract
Whole-exome and whole-genome sequencing technology has enabled the discovery of rare genetic variants associated with human health and diseases. However, existing statistical methods used for rare variant association testing are not well-suited for building genetic risk prediction models that jointly incorporate rare and common variants. We propose STELLAR, a flexible ensemble learning-based approach to compute rare variant polygenic risk scores (PRS) using association summary statistics to enhance conventional common variant PRS. Our method combines burden-based and penalty-based rare variant analysis and leverages functional annotation information to prioritize potentially causal variants within the prediction models. In simulation studies, PRS using STELLAR consistently showed the highest prediction accuracy compared to models using common variants alone or rare variant burdens. Applied to UK Biobank whole-exome sequencing data (n=310,831) across eight continuous and five binary traits, STELLAR significantly improved prediction accuracy, refined stratification of individuals at the highest genetic risk beyond common variants, and prioritized biologically relevant genes. STELLAR provides a scalable strategy to incorporate rare variants into PRS in addition to common variants, advancing precision risk prediction and enabling more comprehensive assessment of genetic contributions to complex diseases.
Krueger, C. J.; Fischer, M.; Rizwan, T.; Kumar, M. M.; Bhargava, S.; Gerszten, R. E.; Taylor, K. D.; Cho, M. H.; Rotter, J. I.; NHLBI TOPMed Consortium, ; Perera, M.; Hu, X.; Manichaikul, A. W.; Im, H. K.; Wheeler, H. E.
Show abstract
Proteomic predictive models are predominantly trained on cis-acting variants in European-ancestry cohorts, limiting power and predictive accuracy in ancestrally diverse populations. We performed cis- and trans-protein quantitative trait locus (pQTL) mapping and developed protein-prediction models using whole-genome sequencing (WGS) and plasma protein levels (Olink) across four ancestry groups from the Trans-omics for Precision Medicine (TOPMed) Multi-Ethnic Study of Atherosclerosis (MESA): European (EUR, n=1270), African (AFR, n=675), Hispanic (HIS, n=642), and Chinese (CHN, n=366), and a combined population (ALL, n=2953). African-ancestry samples demonstrated improved fine-mapping resolution relative to cohort size, yielding significantly smaller cis-credible sets than European-ancestry samples, consistent with shorter linkage disequilibrium (LD) blocks and greater allele frequency diversity in African-ancestry populations. For the first time, we benchmarked fine-mapping models SuSiE, SuShiE, MultiSuSiE, and SuSiEx with multi-ancestral cohorts, revealing a precision-recall tradeoff driven by model assumptions. Comparing protein-prediction models, multivariate adaptive shrinkage (MASHR) and ultimate deconvolution in R (UDR) outperformed elastic net (EN) regression, with trans-pQTL inclusion and fine-mapping improving prediction performance and proteome-wide association study (PWAS) discovery. Applying our models in PWAS of 10 phenotypes, we discovered 68 protein-phenotype associations in All of Us (AoU) that also replicated in Pan-UK Biobank. MASHR and UDR models identified 60% more protein-phenotype associations than EN. Notably, 32 of these associations were not previously reported in the GWAS Catalog. Overall, our study demonstrates the importance of including multiple ancestries in genomic studies to capture the full spectrum of regulatory variation and improve cross-ancestry generalizability.
Tan, T.; Samee, M. A. H.
Show abstract
Genome-wide association studies (GWAS) have identified numerous variant-trait associations; yet, assigning effector genes to GWAS loci remains challenging. Similarity-based machine-learning methods, such as PoPS, prioritize effector genes from shared functional profiles among trait-relevant genes. These models assign a prioritization score for each gene and nominate a single effector gene within a GWAS locus. However, the scores provide limited insight into why a gene was prioritized or whether the nomination is biologically plausible. To address this gap, we introduce Kernelized Polygenic Priority Score, K-PoPS, a kernelized reformulation of PoPS that enables gene-centric explanations by decomposing each prediction into contributions from training genes. For each prioritized gene, K-PoPS reports top contributor genes and an anchor score that quantifies support from a user-defined set of trait-relevant genes. Across 38 Pan-UK Biobank traits, the full-feature OLS implementation underlying K-PoPS improved closest-gene enrichment relative to default PoPS for 26 of 37 evaluable traits. Across 25 traits with curated anchor sets, predictions supported by anchor scores were more enriched for closest-gene proxies than unsupported predictions. When applying to blood level apolipoprotein B, K-PoPS nominated SCARB1 over UBC gene, and further provided convincing explanations that support this prediction. Using explanation evidence, K-PoPS identified multiple plausible effector genes within a dilated cardiomyopathy locus, contrary to the parsimonious assumption. In summary, K-PoPS provides a post hoc framework for examining and interpreting GWAS effector-gene nominations.
Liu, L.; Wang, C.; Kravets, O.; Fermin, D.; Eichinger, F.; Zanoni, F.; Khan, A.; Zhang, J. Y.; Ouyang, Y.; Li, Q.; Hamilton, P.; Kalra, P. A.; Chinnadurai, R.; Reidy, K.; Kopp, J.; Mucha, K.; Smith, C.; Smith, A.; Mcnulty, M.; Eddy, S.; Nair, V.; Helmuth, M.; Klunder, B.; Vasylyeva, T.; Smoyer, W.; Berthier, C.; Parekh, R.; Wenderfer, S.; Martin, T.; Solkolva, K.; Sealfon, R.; Theesfeld, C.; Parsa, A.; Gbadegesin, R.; Sampson, M.; Sanna-Cherchi, S.; Troyanskaya, O.; Paul, D. S.; Petrovski, S.; Goldstein, D.; Mariani, L. H.; Gharavi, A.; Kretzler, M.; Kiryluk, K.
Show abstract
IgA nephropathy (IgAN), IgA vasculitis (IgAV), focal segmental glomerulosclerosis (FSGS), membranous nephropathy (MN), and minimal change disease (MCD) account for the majority of idiopathic glomerulo-nephropathies (GN). These disorders involve immune system dysregulation and have a complex genetic architecture. Currently, there are no adequately powered blood transcriptomic datasets coupled to genetic data from patients with GN that can delineate disease-context specific genetic effects on blood immune cell transcriptome. We performed whole genome sequencing coupled with bulk blood transcriptome sequencing on 1,822 participants from the CureGN study, a prospective cohort of participants with a kidney biopsy diagnosis of primary GN. We generated disease-context specific transcriptome-wide maps of gene expression QTL (eQTL), splicing QTL (sQTL), and double strand RNA-editing QTL (edQTL) for FSGS (N=447), IgAN (N=403), IgAV (N=123), MCD (N=408), and MN (N=441), as well as cross-disease maps for all 1,822 participants. Our QTL mapping identified 16,068 eGenes, 4,644 sGenes and 4,611 edQTLs with an FDR<0.05 in at least one GN type. Approximately 5-10% of the QTL signals were unique to a specific GN type, while ~90% were shared between at least two conditions. Colocalization analysis demonstrated that ~80% of shared eGenes between traits also shared the same causal variants, whereas ~2% had distinct causal variants, suggesting context-specific regulatory effects. Cross-phenotype QTL mapping uncovered 6,466 eGenes, 2,705 sGenes and 5,321 edQTLs not previously detected in GTEx. Age, eGFR, and proteinuria-interaction QTL analyses identified hundreds of loci modified by age and disease severity. Lastly, integrative analyses with GWAS nominated new candidate genes for each of the five GN types under study. In summary, we generated comprehensive maps of GN-context-specific genetic effects on blood transcriptome, providing a powerful resource for integrative gene discovery studies of primary GN.
Takahashi, H.; Hatano, H.; Kono, M.; Haruta, K.; Nakano, M.; Bagherzadeh, R.; Drees, M. M.; Oguma, Y.; Harita, D.; Kawashima, T.; Arakawa, T.; Inokuchi, H.; Nishino, T.; Asahara, K.; Itamiya, T.; Inamo, J.; Natsumoto, B.; Tsuchida, Y.; Sumitomo, S.; Suzuki, A.; Kochi, Y.; Fujio, K.; Yamamoto, K.; Ohta, T.; Kawakami, E.; Ishigaki, K.
Show abstract
Causal variants of complex traits are enriched at transcription factor (TF) binding sites and are thought to contribute to pathology by disrupting TF activity and thereby causing transcriptome dysregulation. However, existing approaches typically address TF-mediated gene regulatory networks (TF-GRNs) and transcriptomes separately, and methods that jointly leverage both to systematically assess disease heritability remain limited. We aimed to develop a framework that jointly leverages TF-GRNs and transcriptomes to assess disease heritability. Here, we constructed a matrix encoding TF-GRNs and developed an unsupervised analytical pipeline, Canonical correlation Analysis of Transcriptome and TF-gene regulatory Networks (CATaN). CATaN applies canonical correlation analysis (CCA) to extract canonical correlation (CC) components, i.e., shared variation components between transcriptomes and TF-GRNs, and converts them into genome-wide functional annotation scores connected to stratified LD score regression (S-LDSC) for heritability analysis. We applied CATaN to eight datasets, including 19,198 bulk samples and 611,772 single cells from human and mouse sources, identifying 588 CC components that are significantly enriched for SNP heritability across 69 complex traits. Notably, functional annotation tracks based on these TF-GRNs are distinct from transcriptome signatures prioritized by LDSC-SEG, with greater heritability enrichment for a subset of traits. Finally, we suggest that CATaN may help prioritize candidate causal variants for experimental fine-mapping using genome editing. Together, integrating TF-GRNs with transcriptomes reveals disease-relevant regulatory programs that are not fully captured by transcriptome-based analyses alone.
Williamson, A.; Carrasco Zanini, J.; Zoodsma, M.; Koprulu, M.; Zaidi, A.; Hunt, K. A.; Taylor-Brill, E. S.; Kohleick, L.; Genes & Health Industry Consortium 1, ; Genes & Health Research Team, ; Chinnery, P. F.; Newman, W.; Finer, S.; Pietzner, M.; van Heel, D. A.; Langenberg, C.
Show abstract
Proteogenomic studies have transformed the way we derive novel insights into human biology and pathophysiology but are currently limited by their proteomic coverage and ancestral representation. Here, we integrate rare and common genetic variation with measurements of >11,000 plasma proteins on two affinity-based platforms (>6,000 targets not previously covered) in 1,535 individuals of British Bangladeshi and Pakistani ancestry of the Genes & Health cohort. We report 3,826 high-confidence common (minor allele frequency (MAF)>1.0%) protein quantitative trait loci (pQTLs), over half of which are novel and including >200 pQTLs with greater MAF in South Asians. Systematic analyses of rare (MAF<1.0%) exonic variants identify 230 gene-protein pairs and highlight the joint and distinct contributions of rare and common variants to inherited differences in protein levels. We expand analyses beyond the nuclear genome and identify 3 mitochondrial pQTLs, including a common variant in MT-RNR1, associated with lower myelin protein zero (MPZ), identifying a potential novel mechanistic link for MT-RNR1's poorly understood role in hearing loss. We create the first proteogenomic disease network in individuals of South Asian ancestry based on 384 cis-pQTL with a shared genetic disease or risk factor signal, including conditions substantially more common in South Asians, such as metabolic diseases or pregnancy-related conditions, providing insights into the underlying mechanisms. In summary, our study demonstrates the value and scientific efficiency of proteomic studies in genetically informative and understudied populations for identifying novel causes of globally relevant diseases.
Narisu, N.; Li, H. X.; Rathbun, C. J. M.; Varshney, A.; Swift, A. J.; Yan, T.; Sinha, N.; Currin, K. W.; Xue, D.; Robertson, C. C.; Taylor, D. L.; Taylor, H. J.; Beck, A.; Lee, B. N.; Wang, L.; Broadaway, K. A.; Wilson, E. P.; Stringham, H.; Saramies, J.; Lakka, T. A.; Spracklen, C. N.; Scott, L. J.; Stitzel, M. L.; Tuomilehto, J.; Laakso, M.; Koistinen, H. A.; Boehnke, M.; Arda, H. E.; Chen, S.; Biesecker, L. G.; Bonnycastle, L. L.; Erdos, M. R.; Mohlke, K. L.; Parker, S. C. J.; Collins, F. S.
Show abstract
Genome-wide association studies (GWAS) have identified >1,200 signals associated with type 2 diabetes (T2D), yet identifying functional variants remains challenging because the majority of them lie in noncoding regions of the genome and are in areas of high linkage disequilibrium (LD). While chromatin accessibility QTL (caQTL) and expression QTL (eQTL) analyses are useful for nominating regulatory mechanisms underlying GWAS signals, limitations still exist in pinpointing functional variants within regions of high LD. A complementary approach that has been less frequently applied is to focus on the allele-specific effect on chromatin accessibility at heterozygous single-nucleotide polymorphisms (SNPs), hereafter referred to as allelic imbalance. We analyzed the allelic imbalance of reads generated from an assay for transposase-accessible chromatin with sequencing (ATAC-seq) across genotyped samples from 490 donors in T2D-relevant tissues: skeletal muscle, liver, pancreatic islets, adipose tissue, and relevant cell types. We identified 119,949 allelically imbalanced SNPs (FDR<0.05) across the genome. The allelic imbalance was often most prominent in one tissue and showed an enrichment overlapping with tissue-specific transcription factor (TF) binding footprints. Focusing on the 8,581 SNPs in previously published 99% credible sets from 338 T2D GWAS signals, we identified 256 imbalanced SNPs across 123 (36.4% of) signals, each showing allelic imbalance in at least one tissue or cell type. Of these, 71 signals contained only a single imbalanced SNP, representing excellent candidate causative variants. As a proof-of-concept, we showed that 23 of the 256 imbalanced SNPs were supported by allelic assays from previous studies. Further, we experimentally validated two imbalanced SNPs as likely functional variants: rs34584161 among a seven-SNP T2D credible set at the RNF6 signal in islets and rs849134 among a 13-SNP credible set at the JAZF1 signal in liver. This study demonstrates the power of integrating ATAC-seq allelic imbalance (ASAI) with GWAS statistical fine-mapping to identify candidate functional regulatory variants from among tightly linked GWAS variants in disease-relevant tissues. While applied here in T2D, this approach represents a widely applicable high-throughput framework for refining the genetic architecture of complex traits.
Vu, H. T. H.; Sun, H.; Kudtarkar, P.; Sharp, S. A.; Brusman, L.; Wang, Y.; Huang, Y.; Mao, R.; Feng, F.; Corban, S.; Huber, A. K.; Shilin, A.; Sun, Y.; Narayanaswamy, S.; Jang, D.; Jurgens, J.; Robertson, C. C.; Shrestha, S.; Bate, T.; Nguyen, T.; Smadbeck, P.; Zhang, L.; Brandes, M.; The PanKbase Consortium, ; Flannick, J.; Burtt, N.; Chen, S.; Liu, J.; Cartailler, J.-P.; Voight, B. F.; Stitzel, M. L.; Brissova, M.; Gloyn, A. L.; Gaulton, K. J.; Parker, S. C. J.
Show abstract
Aims/hypothesisSingle-cell RNA sequencing (scRNA-seq) of pancreatic islet tissue is a powerful tool for investigating Type 1 Diabetes (T1D). However, individual datasets are limited in size and fragmented across donors, laboratories, and experimental conditions, highlighting the need for a unified single-cell atlas. This study aimed to construct a comprehensive, integrated scRNA-seq map of human isolated pancreatic islets by collating data from diverse sources. MethodsPublicly available scRNA-seq datasets derived from isolated pancreatic islets, generated and/or provided by the Human Pancreas Analysis Program (HPAP), Prodo Labs, and the Integrated Islet Distribution Program (IIDP), were collected. Systematic quality controls were implemented to select high-quality samples, reads and cells. Data integration was conducted, accounting for important variables such as age, sex, body mass index (BMI), origin study, treatments, islet data/distribution resources, and sequencing chemistry. ResultsWe generated a comprehensive single-cell atlas of human pancreatic islets comprising 191 high-quality assays from 140 donors (59 female, 81 male) across five phenotypic groups: controls without diabetes (69 donors), autoantibody-positive donors without diabetes (12), pre-diabetic donors (11), donors with type 1 diabetes (12), and donors with type 2 diabetes (36). The atlas also includes experimentally perturbed samples, including those exposed to SARS-CoV-2 infection and pro-inflammatory cytokines. In total, the atlas contains 448,935 cells, capturing major endocrine islet populations, such as alpha cells (43.3%) and beta cells (26.8%), as well as non-endocrine populations such as endothelial cells (0.75%) and immune cells (0.6%). Conclusions/interpretationBy uniformly harmonizing and integrating data from multiple sources, we have developed a comprehensive single-cell atlas of isolated human pancreatic islets, which is publicly available at www.pankbase.org. The atlas provides a platform for hypothesis-driven investigation of diabetes pathophysiology and, given rigorous quality control at the read, barcode, and sample levels alongside careful metadata curation, is well suited for downstream machine-learning applications.
The ENCODE Project Consortium, ; Reddy, T. E.
Show abstract
We present the Encyclopedia of DNA Elements (ENCODE), a reference map of the genomic basis of gene regulation. A product of more than two decades of systematic interrogation of genome function, ENCODE encompasses more than 16,000 genome-wide experiments, predominantly in primary cells and tissues, focused on three core layers of genome function. First, ENCODE now provides a catalog of gene regulatory elements. The catalog is based on a foundation of 5.3 million DNase I hypersensitive sites that delineate essentially all chromatin-accessible regulatory DNA in the human genome, as well as extensive maps of chromatin states, transcription factor occupancy, and nascent transcription, and systematic predictions of the functional consequences of non-coding genetic variants on regulatory element activity. Second, ENCODE expands the catalog of genes and transcripts, which now includes nearly 18,000 novel human long noncoding RNA genes, nearly 150,000 novel transcript isoforms, and genome-wide maps of transcript stability across cell types and time. Third, ENCODE now maps physical and functional interactions among regulatory elements and genes across more than 100 human tissues and cell lines at up to 10 bp resolution. Those studies reveal a vast network of interactions among millions of loop anchors across and links those interactions to gene expression. Through parallel studies in mice, ENCODE also provides extensive maps of gene regulatory elements, transcripts, and their interactions across the mouse postnatal development. Together, the Encyclopedia of DNA Elements provides a foundational framework for genome-focused studies of human and mouse biology.
Huang, N.; Ragsac, M. F.; Gui, X.; Tantisira, K. G.; Amariuta, T.
Show abstract
Asthma is a heritable complex disease that disproportionately burdens minority and admixed populations in the US. However, the causal genes and regulatory mechanisms governing inherited risk remain largely unresolved. We performed a European-ancestry meta-analysis of 141,894 cases and 1,361,846 controls drawn from the Trans-national Asthma Genetic Consortium (TAGC) and Global Biobank Meta-analysis Initiative (GBMI), yielding an estimated h2SNP of 0.056 (SE = 0.0038) and 275 independently associated loci. To enhance mechanistic inference beyond variant-level associations, we developed a multimodal framework to predict asthma risk integrating GWAS summary statistics, bulk tissue expression quantitative trait loci (eQTL) data from the Genotype-Tissue Expression (GTEx) project, and single-cell gene eQTL data from the OneK1K Project. We performed transcriptome-wide association studies (TWAS) and subsequently applied probabilistic fine-mapping with FOCUS to prioritize putative causal genes expressed in bulk tissues and higher resolution immune cell populations. Fine-mapping asthma-associated genes implicated barrier-immune and metabolic-endocrine tissues alongside adaptive T-cell subsets as the primary mediators of asthma genetic risk, resolving canonical CD4+ Th2 effector genes including IL1RL1, TSLP, STAT6, and GATA3. Using these prioritized genes, we constructed a polygenic transcriptome risk score (PTRS) using random forest to integrate gene-level effects across critical tissues and cell types. Evaluated in two ancestrally distinct pediatric asthma cohorts, the Childhood Asthma Management Program (CAMP) and the Genetics of Asthma in Costa Rica Study (GACRS), our PTRS demonstrated improved transferability over the standard variant-level and gene-level baseline models. While modest common variant heritability limits the discriminative power of our models, we estimated a theoretical maximum achievable area under the receiver operating characteristic (AUROC) curve of 0.64. Our integrative nonlinear model of PRS-CSx and cross-modal (bulk tissue and single cell) FOCUS PTRS resulted in the best cross-cohort performance (CAMP AUC = 0.632, sd = 0.04, 3.55 case/control odds ratio in top vs. bottom quartiles), representing an increase of +0.118 AUC over PRS-CSx, +0.067 AUC over tissue-specific TWAS pruning and thresholding, and +0.041 AUC over cell-type-specific FOCUS PTRS. Our results demonstrate that modeling nonlinear interactions between variant- and gene-level effects across both bulk tissue and single cell eQTL data improves our ability to determine high-risk individuals and to explain the likely mechanisms driving genetic susceptibility of childhood-onset asthma.
Edgar, R. D.; Portman, J. R.; Hu, H.; Pouyabahar, D.; Rahman, R. R.; Stueckmann, D.; Choi, Y.; Neavin, D. R.; Atif, J.; Clarke, Z. A.; Gao, R.; Khare, S.; Li, Z.; Martens, L.; Murti, A.; Nakib, D.; Shirgaonkar, N.; Thomann, S.; Thone, T.; Wilson-Kanamori, J. R.; Breitkopf-Heinlein, K.; Lattouf, E. I.; Li, R.; Napoliello, R.; Rahbari, N. N.; Sadria, M.; Yakubovsky, O.; Andrews, T.; Aronow, B. J.; Cuenca, A. G.; DePasquale, E. A. K.; Huppert, S. S.; Itzkovitz, S.; Lauer, G. M.; Mysore, K. R.; Powell, J. E.; Schwartz, R. E.; Sharma, A.; Taylor, S. A.; Vallier, L.; Wang, B.; Dasgupta, R.; Grün, D
Show abstract
The human liver is composed of a heterogeneous mix of cell types. How these distinct populations contribute individually and collectively to liver function remains poorly understood. Although single-cell technologies have advanced our understanding of liver biology, individual studies have often been limited by small donor cohorts and inconsistent cell type annotations. Integrating multiple datasets can overcome these challenges and better capture biological variability. We present the Human Liver Cell Atlas (HLiCA), an integrated reference of non-disease liver cells assembled from eight datasets across six research centers, encompassing more than 525,000 cells from 110 donors. Developed in collaboration with the Human Cell Atlas Liver Bionetwork, the HLiCA incorporates expert-curated cell annotations refined through community feedback and dedicated cell type annotation meetings. The HLiCA classifies cells into six lineages and expands the cell type resolution to include 47 distinct cell types. Starting from raw sequencing reads, we realigned all data and performed rigorous benchmarking to ensure robust integration across technical and biological variables. Genetic ancestry was inferred for all samples to evaluate the range of ancestral backgrounds represented in the atlas. The expanded cell type annotation enabled identification of previously unrecognized liver cell types, including NRXN1+ stromal cells. Their presence was validated using spatial transcriptomics, which localized NRXN1+ stromal cells to periportal regions. With the number of donors included in the HLiCA we were able to examine cell type specific associations with demographic covariates. In hepatocytes, drug metabolism genes showed differential expression between sexes, and in cholangiocytes, mucus-production genes varied with age. As the largest and most genetically diverse human liver cell atlas to date, the HLiCA provides a comprehensive, well-annotated reference for the field, annotated by expert consensus. This resource will enable deeper interrogation of liver cellular diversity, architecture, and function in the healthy human liver and serve as a reference to understand changes that occur with disease.
Tian, Y.; Wong, J.; McDonnell, S.; Zhong, H.; Wu, L.; Larson, N.; Manley, B. J.; Wang, L.
Show abstract
Long-read nanopore sequencing enables simultaneous detection of germline variation and native DNA base modifications on individual DNA molecules, providing a unique opportunity to investigate allele-specific epigenetic regulation. Here, we performed whole-genome nanopore sequencing on normal and tumor prostate tissues to characterize differential methylation, methylation entropy, and allele-specific methylation (ASM) associated with noncoding genetic variants. Genome-wide analysis identified extensive cancer-associated differentially methylated regions (DMRs), with hypermethylated DMRs significantly enriched near transcription start sites and transcriptional regulatory regions. Integration with transcriptomic datasets revealed strong inverse relationships between promoter methylation and gene expression, while 5-hydroxymethylcytosine (5hmC) levels positively correlated with transcriptional activity across gene bodies. Using fragment-level methylation patterns enabled by long-read sequencing, we further quantified methylation entropy incorporating both 5mCG and 5hmCG states. Cancer-hypermethylated DMRs exhibited markedly reduced entropy, consistent with clonal fixation of methylation states during tumor progression. Entropy profiling across chromatin annotations demonstrated maximal epigenetic heterogeneity at partially modified enhancer-associated regions. To investigate cis-regulatory genetic effects, we developed a simple ASM framework (nanoASM) that can partition sequencing reads by allelic state and identifies allele-specific DMRs directly from long-read data. Compared with conventional population-level mQTL analysis, ASM demonstrated substantially improved statistical efficiency by leveraging within-individual contrasts and reducing sample-level heterogeneity. Although germline single nucleotide polymorphisms (SNPs) were largely shared between normal and tumor tissues, ASM patterns differed substantially, with tumor-associated ASM regions displaying significantly larger genomic span and stronger allelic methylation differences. Comparative analysis with TCGA prostate mQTL and GTEx prostate eQTL datasets demonstrated substantial concordance between ASM directionality and downstream transcriptional effects, particularly for variants located within DMRs and near transcription start sites. At the IRX4 prostate cancer risk locus, ASM identified an androgen-responsive regulatory domain overlapping AR ChIP-seq and H3K27ac peaks, nominating rs6885084 as a candidate functional variant. At the PSCA locus, ASM anchored by rs4736369 was associated with allele-specific methylation, chromatin activation, transcript abundance, and isoform usage. Together, these findings establish nanopore-based ASM analysis as a powerful approach for resolving functional noncoding variants and their regulatory domains they control in prostate cancer.
Onawole, A.; Adegoke, R. A.; Amoo, O.
Show abstract
Polygenic scores summarise genetic predisposition to a trait, but a population-level accuracy figure cannot tell a clinician whether a given prediction is reliable for the person in front of them. This gap is most consequential for individuals whose ancestry is under-represented in the discovery cohort, precisely the patients for whom a wrong trust call carries the highest clinical cost. We present TrustPGS, a framework that tells clinicians and downstream models which individual predictions can be trusted and which cannot, so that polygenic scores can inform clinical decisions rather than being acted on uniformly regardless of how well-supported each prediction is. The framework rests on two axes calibrated on a discovery cohort, the consensus of a Bayesian posterior-sample ensemble and the directional agreement of the top-magnitude linkage-disequilibrium blocks. We computed SBayesRC posterior-sample scores for ten polygenic traits in the 1000 Genomes Project phase-3 cohort and tested whether the resulting trust labels transfer, without recalibration, to the ancestrally diverse Simons Genome Diversity Project, comparing strict application of the European cutoffs, percentile-rank rescaling, and within-cohort recalibration. Percentile-rank rescaling preserved an enrichment factor above one in non-European populations for five of ten traits (Alzheimer disease, breast cancer, body mass index, LDL cholesterol, and systolic blood pressure), traits whose European and target-cohort distributions were shifted but comparable in shape. Three traits (coronary artery disease, height, and schizophrenia) carried distributions that differed in shape rather than location, a pattern traceable to discovery-cohort bias that recalibration could not repair either, and two further traits (type 2 diabetes and educational attainment) showed intermediate behaviour, present but never enriched in one case, and an apparent success that rank-mapping correctly unmasked as artefactual in the other. Because each of these patterns is detectable before any individual-level claim is made, TrustPGS gives clinicians and downstream models a falsifiable, per-trait basis for deciding when a reliability label can be trusted on a new population, rather than a single portability promise that holds or fails silently.
Ahn, K.; House, J. S.; Burkholder, A.; Tran, T. C.; Breeyear, J. H.; Justice, C. M.; Durney, J.; Jones, A. M.; Reyes, P. S.; Bailey, M. H.; Davis, M. F.; Vicenti, A. T.; Karnes, J. H.; Hollenbach, J. A.; Fargo, D. C.; Ginsburg, G. S.; Woychik, R. P.; Denny, J. C.; Motsinger-Reif, A. A.
Show abstract
The human leukocyte antigen (HLA) region is the strongest genetic contributor to many immune-mediated diseases, yet whether HLA architecture is shared across ancestries remains unclear. We analyzed high-resolution HLA variation in 390,823 participants from the All of Us Research Program spanning six genetic ancestry groups, including 262,915 with linked electronic health records. Using whole-genome sequencing and graph-based inference, we genotyped 20 HLA genes at G-group resolution and identified 4,780 distinct alleles. Analyses accounting for disparate sample sizes demonstrated that ancestry-private allelic variation reflected unequal discovery depth rather than ancestry-population specificity. A meta-analysis of ancestry-stratified phenome-wide association analyses with 363 HLA alleles with frequency > 0.001 and 3,430 clinical phenotypes identified 1,461 significant HLA-phenotype associations (FDR < 0.05). Although many associations reached significance in only one ancestry group, effect directions were largely concordant, highlighting differences in allele frequency, linkage disequilibrium, and statistical power among ancestry groups. Stepwise conditional modeling demonstrated that common complex trait variation could be concurrently explained by five to seven independent HLA allele signals. These findings demonstrate that a multi-ancestry, phenome-wide study can distinguish true biological heterogeneity from sampling-driven detectability differences in HLA.
Hu, S.; Zhu, P.; Wu, S.; Gao, S.; Wang, R.; Liu, F.; He, Y.; Han, Z.; Wang, T.; Wang, M.; Ren, C.; Ji, X.; Zhao, W.; Li, S.; Liu, G.
Show abstract
Transient ischemic attack (TIA) is a critical harbinger of subsequent stroke, and most genetic risk remains uncharacterized. Here we firstly performed the largest TIA genome-wide association study (GWAS) meta analysis in 1,332,453 European individuals (58,976 cases and 1,273,477 controls), followed by an independent replication in 610,409 non-European individuals (23,557 TIA and 586,852 controls), a multi-ancestry GWAS meta analysis in 1,942,862 individuals (82,533 cases and 1,860,329 controls), and a cross-trait GWAS meta-analysis of TIA with stroke and its subtypes. We identified 44 loci including 25 known stroke loci and 19 TIA specific loci (CELSR2, SLC4A7, CASC15, SRRM3, SLC44A1, LOC107984361, GSE1, LOC105372530, HCG20, OXR1, SLC4A1, RBBP8, TUSC3, DCC, PALMD, ZNF475, CTAGE1, FUT2 and MRPS6). Post-GWAS pinpointed 51 high confidence genes (24 are potential therapeutic targets) and 13 statistically significant pathways including protein-lipid complex, neurofibrillary tangle, high-density lipoprotein particle. These findings provide critical insights into the genetic basis of TIA.
Wu, D.; Yang, C.; Chen, Q.; Suo, M.; Zhou, F.; Liu, A.; Yu, D.; Nie, L.; Yang, T.; Sun, Y.; Han, J.; Yang, L.; Ni, Q.; Sun, D.; Lu, Y.; Fu, L.; Yang, Y.; Yu, J.; Qi, J.; Dai, W.; Yang, X.; Qiu, L.; Yang, D.; Jiao, Y.; Zhou, F.; Zhang, W.; Wang, F.; Yang, Y.; Zeng, Z.; Feng, Z.; Chen, Y.; Li, Y.; Li, Y.; Zhao, S.; Long, A.; Wang, Z.; Li, Q.; Zhao, R.; Ding, G.; Wang, Q.; Tuo, Y.; Yu, J.; Li, H.; Liu, K.; Zhang, Y.; Yan, X.; Dawa, D.; Zhang, Y.; Bi, A.; Chen, G.; Qian, S. H.; Li, X.; Bi, X.; Liu, J.; Li, J.; Fu, K.; Ye, S.; Wang, S.; Yang, J.; Zhou, Q.; Jiang, J.; Xu, W.; Liu, Y.; Liu, A.; Meng,
Show abstract
East Asian populations, representing over 20% of the global population, remain critically underrepresented in human genomic studies, limiting our understanding of population-stratified genetic variation and its implications for health and disease. Here we present the first phase of the Asian Pan-Genome project (APG), comprising 320 nearly complete, fully phased haploid genome assemblies from 160 East Asian individuals. These assemblies achieve unprecedented quality, with an average contig N50 of 144.3 megabase pairs and an average quality value of 64.5. Leveraging these superior assemblies, we reveal previously uncharacterized diversity in human repeatome, including population-stratified patterns in centromere satellites and rDNA arrays. Compared to existing global human genome assemblies, the newly generated genomes supplement 152 million base pairs of novel sequences, 355 gene gains, 18,300 structural variation loci and 26 large euchromatic inversions missing from current human pangenomes. We perform population stratification analyses of structural variations, and further resolve the structural haplotypes of complex genomic regions such as Major Histocompatibility Complex and Survival Motor Neuron loci across global pangenomes, exemplifying tandem-duplicate and inversion-rich complex locus architectures in the human genome, respectively. This resource provides a critical foundation for human genetic studies, especially for East Asian populations, promoting more accurate variant discovery, reducing bias, and ultimately advancing the equity and efficacy of genomic medicine.